Speaker diarization is a key component for multiple downstream speech technologies, including speech transcription, meeting analytics, and conversational understanding; however, Romanian lacks publicly established diarization resources and benchmarks. This paper evaluates cross-lingual transfer of diarization systems pretrained on predominantly English data, under a strict no-adaptation policy. We compare an end-to-end neural diarization approach (MSDD) and a traditional modular pipeline (segmentation + speaker embeddings + clustering), both used as is with pretrained components. To enable controlled analysis despite the lack of Romanian diarization datasets, we construct a synthetic Romanian conversational benchmark with explicit conditions on speaker count (2–5) and overlap regime (no overlap versus overlap). We report the diarization error rate (DER) and Jaccard error rate (JER) across all conditions, analyze sensitivity to overlap and the number of speakers, and provide an error-component breakdown to identify dominant failure modes. Across all conditions, the end-to-end system outperforms the pipeline (DER 0.140 versus 0.267; JER 0.152 versus 0.320). Performance degrades with overlap and with increasing speaker count in both paradigms, with speaker confusion dominating the additional error under overlap.
Loading....